Papers with Tatoeba database

1 papers
TaPaCo: A Corpus of Sentential Paraphrases for 73 Languages (2020.lrec-1)

Copied to clipboard

Challenge: a crowdsourcing project aimed at language learners has created a paraphrase corpus for 73 languages . the corpus contains 1.9 million sentences, with 200 - 250 000 sentences per language .
Approach: They propose to use a Tatoeba-based dataset to create a paraphrase corpus for 73 languages.
Outcome: The proposed dataset contains 1.9 million sentences and 200 - 250 000 sentences per language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations